Back

Genetics Selection Evolution

Springer Science and Business Media LLC

Preprints posted in the last 90 days, ranked by how well they match Genetics Selection Evolution's content profile, based on 39 papers previously published here. The average preprint has a 0.02% match score for this journal, so anything above that is already an above-average fit.

1
Optimizing genomic selection: A comparison of SNP selection strategies for reduced-density panels in beef cattle

Ogunbawo, A. R.; Mulim, H. A.; Hidalgo, J.; Ventura, H. T.; Souza, N. O.; Oliveira, H. R.

2026-08-31 genetics 10.64898/2026.08.26.747408 medRxiv
Top 0.1%
61.7%
Show abstract

The exponential increase in the number of genotyped animals, combined with the availability of high-density SNP chips has introduced computational challenges for routine genomic evaluations, particularly during the construction of the genomic relationship matrix. Although higher-density SNP panels can facilitate the identification of causal mutations, their use substantially increases computational requirements without a proportional gain in genomic prediction performance. To optimize computational efficiency while maintaining accuracy of genomic predictions, this study compared five SNP selection strategies (i.e., random sampling, random sampling with inclusion of informative SNPs, linkage disequilibrium (LD)-based pruning, a Shannon entropy-based machine learning approach, and [[EQUATION]]-based prioritization) to develop reduced-density panels for Nellore cattle. Using high-density (HD) genotype data comprising 437,650 SNPs from 304,782 animals (after quality control) as reference, three reduced-density panels (25K, 45K, and 65K SNPs) panels were tested across five traits (i.e., Age at first calving, Stayability, Weaning weight, Yearling weight, Muscling) with diverse genetic architectures. Genomic estimated breeding values (GEBVs) derived from these reduced panels were compared to those obtained from the HD reference panel using Pearsons correlations, under both genomic best linear unbiased prediction (GBLUP) and single-step GBLUP (ssGBLUP) methods. In the GBLUP model, prediction accuracy generally improved with increased marker density. Random selection with and without the informative SNPs consistently yielded the highest accuracies, whereas the [[EQUATION]]-based approach showed the lowest agreement with the HD reference across all densities. In contrast, ssGBLUP demonstrated strong robustness to marker reduction, producing uniformly high correlations {approx}1.00) across all SNP densities and selection strategies. These findings indicate that optimized low-density SNP panels maintain prediction accuracy comparable to HD panels, offering a cost-effective tool for large-scale genomic evaluations.

2
Genomic Inbreeding and Selection Signatures analyses in the Doberman Pinscher breed

Mulim-McCarthy, H.; Fragomeni, B.; Liu, S.; Rojas de Oliveira, H.

2026-06-08 genomics 10.64898/2026.06.04.730131 medRxiv
Top 0.1%
18.7%
Show abstract

The Doberman Pinscher population has undergone strong artificial selection for morphology and behavior, which can reduce genomic diversity and increase autozygosity. Here, we characterized the genome structure and identified selection signatures in Doberman Pinschers using complementary within- and between-population approaches. Genotypes from 3,226 Dobermans Dogs (Illumina CanineHD; 216,184 SNPs) provided by the Doberman Diversity Project were analyzed after purpose-specific quality control. Genomic inbreeding was quantified using four allele-frequency-based metrics and the runs of homozygosity (FROH) approach. Selection signatures were detected using intrapopulation (i.e., Runs of Homozygosity--ROH; Integrated Haplotype Score--iHS; and Number of Segregating Sites by Length--nSL) and interpopulation methods (i.e., Fixation Index--FST; Cross-Population Extended Haplotype Homozygosity--XP-EHH; and Cross-Population Number of Segregating Sites by Length--XP-nSL) comparing the Doberman Pinscher breed to Labrador Retriever (n=237). Dobermans showed high overall inbreeding, with a mean FROH of 0.42 (range 0.22-0.68), whereas the allele-frequency-based inbreeding estimators had similar means ([~]0.04). The partitioning of the ROH indicated high contributions from medium-to-long ROHs, consistent with recent inbreeding. The ROH scans identified 39,512 SNPs in ROH islands ([≥]50% frequency across individuals), with notable concentrations on CFA2, CFA3, and CFA31. Haplotype-based scans identified 2,820 candidate iHS SNPs and 2,173 candidate nSL SNPs (|score|>2). A common set of 310 SNPs was shared among ROH, iHS, and nSL, mapping near 279 genes that were mostly enriched for developmental pathways, particularly neurodevelopment and neuron-related cellular components. Between breeds, 349 highly differentiated SNPs were detected by FST, while XP-EHH and XP-nSL highlighted over 1,000 of Doberman-specific haplotype signals. A total of seven SNPs overlapped across FST, XP-EHH, and XP-nSL, which were located mainly on CFA8 ([~]59.48-60.61 Mb) near the KCNK10, SPATA7, PTPN21, NEGR1, and BTG1 genes. These genes are mainly linked to neural development and signaling, but BTG1 has also been associated with cardiomyocyte cell-cycle regulation, and KCNK10 with cardiac excitability and remodeling. Overall, the Doberman Pinscher breed exhibits high genome-wide autozygosity and levels of inbreeding. In addition, our results showed consistent, multi-method evidence of selection at loci associated with neurodevelopmental and regulatory pathways. These findings provide prioritized targets for follow-up studies that integrate phenotypes relevant to breed health and performance.

3
Variant intolerance scores in cattle

Lanigan, S.; Derks, M. F.; Johansson, A. M.; Johnsson, M.

2026-07-30 genetics 10.64898/2026.07.28.741161 medRxiv
Top 0.1%
18.4%
Show abstract

Variant intolerance methods score the essentiality of genes based on large datasets of genetic variants and have been used in population genomics of humans and model organisms. In this paper, we estimated Residual Variation Intolerance Scores for protein-coding genes and predicted protein domains in cattle. In agreement with results from other species, the most variant-tolerant genes and domains included genes related to olfaction and adaptive immunity, whereas the least-variant tolerant genes and domains included genes involved in fundamental cellular processes. There was a moderate positive correlation with estimates from orthologous human genes. We provide estimates of variant intolerance for cattle may be useful for genomic analyses of deleterious variants and population genomics in cattle.

4
Genetic Modeling of Dyadic Behavioral Traits: Implications for Estimation and Interpretation of Variance Components

Jiang, X.; Siegford, J.; Steibel, J. P.

2026-06-12 genetics 10.64898/2026.06.10.731434 medRxiv
Top 0.1%
12.6%
Show abstract

Studying the genomic control of dyadic social interactions is gaining traction in animal genetics. However, genetic modeling of social interactions poses several challenges, one of which is whether social interactions should be treated as dyadic traits or as aggregated traits at the individual level. In this study, we systematically compared two approaches: dyadic models using dyadic traits and marginal models using marginally aggregated traits and we derived the algebraic relationships between their variance components. In the application, we used a published dataset on post-mixing aggression in pigs, including both directed and undirected aggression records collected during the 9-hour period after mixing among 797 finishing pigs in 59 social groups, as an example to show how model choice can affect variance estimation. Results showed that dyadic models can estimate genetic effects and permanent environmental effects by exploiting repeated dyadic interaction records, thereby enabling a more complete understanding of the sources of variation underlying social interactions. In contrast, marginal models can bias the estimation and interpretation of genetic components, as the aggregated genetic variance may be confounded with other variance components due to the aggregation of dyadic traits. Marginal models may also lead to overestimation of social group and residual variance. These results can provide useful guidance for choosing appropriate modeling strategies for social interaction traits.

5
An endogenous retrovirus insertion disrupting bovine ALKBH8 causes a failure-to-thrive syndrome with immunodeficiency associated with juvenile mortality in Brown Swiss cattle

Glatthard, S.; Kadri, N. K.; Seefried, F. R.; Voitl, L. R.; Weber, B. A.; Schwarzenbacher, H.; Meister, S. L.; Gurtner, C.; OGrady, J. F.; Osbahr, M.; Leonard, A. S.; Meylan, M.; Pausch, H.; Droegemueller, C.; Jacinto, J.

2026-07-10 genomics 10.64898/2026.07.09.737535 medRxiv
Top 0.1%
10.1%
Show abstract

The Brown Swiss (BS) cattle breed is one of the major Swiss dairy breeds. Intensive selection and the widespread use of few elite sires in artificial insemination have increased inbreeding and the occurrence of deleterious recessive alleles in the homozygous state. Analyzing life trajectories in large, genotyped cohorts can identify hidden recessive disorders that are difficult to detect using traditional case-control association testing. Long-read DNA sequencing enables precise detection of causal alleles, including structural variants. This study aimed to (1) identify cryptic recessive loci affecting rearing performance in Swiss BS cattle, (2) evaluate their impact on survival, (3) characterize the associated phenotype, (4) identify the causal variant using long-read whole-genome sequencing, and (5) assess its functional impact. Using Homozygous Haplotype Enrichment/Depletion (HHED) mapping, we identified a risk haplotype (BH39) on chromosome 15 spanning from 16,276,819 bp to 16,446,984 bp that was associated with increased juvenile mortality within the first 180 days of life when present in the homozygous state. The BH39 occurred at a frequency of approximately 4.5% in Swiss BS cattle and 5.3% in German and Austrian BS cattle, and homozygous carriers exhibited a significantly reduced first-year survival rate. Five females homozygous for BH39 underwent clinical examination. They all showed recurrent respiratory disease, impaired growth, poor body condition, rough hair coat, and brown-discolored teeth. Pathological examination revealed bronchopneumonia and eosinophilic enteritis. Clinicopathological findings indicated failure to thrive and immunodeficiency. Long-read WGS of two BH39 homozygous calves revealed a private homozygous coding variant that was in high linkage disequilibrium with BH39. The identified structural variant was an insertion of a large transposable element (10.4 kb ERVK[2-1-LTR]) into the third exon of ALKBH8 (NM_001080341.2 c.267_268indel). Full-length RNA sequencing of cerebellum and liver from a homozygous calf revealed that the endogenous retrovirus (ERV) insertion introduces a cryptic transcription termination signal, truncating ALKBH8 mRNA. This study demonstrates that exploring population-scale genomic data and mining thousands of life-history records, followed by veterinary follow-up evaluations and molecular genetic analyses, provides an effective strategy for identifying cryptic recessive disorders that shorten the lifespan of cattle. The findings provide strong evidence that the ERV insertion into the coding sequence of ALKBH8 represents a loss-of-function variant that causes a previously undescribed recessive disorder that results in increased rearing loss. Interpretive summaryWe identified a recessive disorder in Brown Swiss cattle that causes retarded growth, recurrent infections, immunodeficiency, and increased mortality during the first year of life. Using population-scale genomic data, clinical investigations, and long-read sequencing, we linked the disorder to an exonic transposable element insertion disrupting ALKBH8. The identification of the causal variant now enables direct genetic testing and the implementation of genome-based mating strategies to avoid carrier-by-carrier matings and, consequently, prevent the birth of affected homozygous offspring. We demonstrate the utility of integrating large-scale breeding records, veterinary phenotyping, and advanced genomics to identify hidden defects affecting livestock health and productivity.

6
Genomic insights into bacterial kidney disease resistance in Arctic charr (Salvelinus alpinus) via a 72k SNP array

Palaiokostas, C.; Jeuthe, H.; Nilsson, K. N.; Hallbom, H.; Axen, C.; Evensen, O.; Eriksson, S.; Johnsson, M.

2026-06-27 genetics 10.64898/2026.06.25.734482 medRxiv
Top 0.1%
7.9%
Show abstract

Selection for disease resistance forms one of the most highlighted areas of aquaculture breeding. A breeding program for Arctic charr has been operating in Sweden for over 40 years, making it the oldest of its kind worldwide for this species. However, the lack of available genomic resources prevented selection for any disease-resistance traits. A 72k Axiom SNP array was produced in this study and used to assess the potential to select for charr resistant to bacterial kidney disease (BKD), which is currently a major threat to the industry. Following a challenge experiment with Renibacterium salmoninarum, the causative agent of BKD, relevant phenotypic proxies were collected from approximately 2,000 charr. Thereafter, those animals were genotyped with the new 72k SNP array. The magnitude of the estimated variance components suggested potential for breeding for BKD resistance in charr, with relevant heritabilities ranging from 0.05 to 0.56 depending on the resistance proxy used. In addition, GWAS suggested that BKD resistance is a polygenic trait. Furthermore, genomic prediction approaches indicated that BKD-resistant animals can be identified using their SNP genotypes. Accuracies, expressed as Pearson correlation coefficients, when BKD resistance was analysed as a continuous trait, ranged from 0.42 to 0.52. In the scenario where BKD resistance was treated as a binary trait, the efficiency of genomic prediction was assessed using ROC curves, with an area under the curve of 0.72. Finally, no unfavourable correlations were found with growth traits. The developed 72k SNP array has the potential of being a pivotal tool for the Swedish Arctic charr breeding program. Moreover, our data support the use of genomic prediction in breeding BKD-resistant Arctic charr. As a critical next step, further validations in actual industry conditions would be required.

7
Toward routine health phenotyping: High-throughput prediction of metabolic, immune, and inflammatory biomarkers from milk mid-infrared spectroscopy in early-lactation dairy cows

Ho, P.; Hemsworth, J.; Reich, C.; Bath, C.; Liu, Z.; Rochfort, S.; Khansefid, M.; Tahir, S.; HaileMariam, M.; Goddard, M. E.; Marett, L.; Williams, R.; Ho, C.; Berkhout, M.; Xiang, R.; Chamberlain, A.

2026-07-20 genetics 10.64898/2026.07.14.738588 medRxiv
Top 0.1%
5.0%
Show abstract

This study evaluated the potential of milk mid-infrared (MIR) spectroscopy, combined with routinely available on-farm variables, for predicting serum metabolic, immune, and inflammatory biomarkers in early-lactation cows. Data included 5,936 blood samples from 4,442 cows across 23 Australian dairy herds, with paired milk MIR spectra and serum measurements for up to 14 biomarkers. Prediction models were developed using partial least squares regression and evaluated using nested 10-fold random cross-validation and leave-one-herd-out validation. The results show that while basic herd-test data, including milk fat, protein, and lactose concentration, as well as on-farm variables, including DIM, calving age, breed, and herd could predict serum biomarkers, combining MIR spectra with these on-farm variables produced the best overall performance. In random cross-validation, blood urea nitrogen (BUN) was predicted most accurately (R2 = 0.78), while {beta}-hydroxybutyrate (BHB) and nonesterified fatty acids (NEFA) showed moderate accuracy (R2 = 0.56 and 0.44, respectively). BUN also showed the strongest external validation performance, with leave-one-herd-out R2 = 0.58 and comparable accuracy for predicting records collected after 70 days in milk (R2 = 0.65). BHB and NEFA had moderate leave-one-herd-out accuracy but did not transfer beyond early lactation. Most other biomarkers showed low or inconsistent external validation performance. Overall, MIR spectroscopy combined with on-farm variables shows promise for routine prediction of BUN, BHB and NEFA, which can be used for monitoring and genetic evaluation of, for example, ketosis and energy deficit. Initial random cross-validation results for glucose, bilirubin and cholesterol were promising, but more data is needed to improve the prediction accuracy and robustness of the predictions. HighlightsO_LIMilk MIR can predict several serum biomarkers in early-lactation dairy cows. C_LIO_LIMIR-predicted blood urea nitrogen shows the greatest accuracy and robustness. C_LIO_LIMIR-predicted {beta}-hydroxybutyrate and nonesterified fatty acids show moderate accuracy. C_LIO_LIMost mineral, hepatic, and inflammatory biomarkers had limited accuracy. C_LIO_LIRoutine MIR phenotyping is most promising for BUN, BHB, and NEFA. C_LI SummaryMilk mid-infrared (MIR) spectroscopy and on-farm variables are evaluated as a high-throughput tool to predict health-related serum biomarkers in early-lactation dairy cows. The data include 5,936 paired blood and milk samples from 4,442 cows across 23 Australian dairy herds and up to 14 serum biomarkers. Prediction models are developed using partial least squares regression with nested random cross-validation and leave-one-herd-out validation. Blood urea nitrogen (BUN) shows the greatest and most transferable prediction accuracy across herds and lactation stages. {beta}-hydroxybutyrate (BHB) and nonesterified fatty acids (NEFA) are predicted with moderate accuracy, but only during early lactation. Random cross-validation results for glucose and bilirubin are promising, but larger datasets are needed for robust external validation. Most other mineral, hepatic, and inflammatory biomarkers show limited external prediction accuracy. These results indicate that MIR-based routine health phenotyping is most promising for BUN, BHB, and NEFA.

8
A cis-regulatory variant in ASIP causes gray coat color in the donkey

li, y.; Liu, Y.; wu, j.; liu, s.; lin, x.; guo, k.; yang, t.; feng, m.; zhang, h.; wang, x.; xing, w.; qian, s.; yang, r.; zhao, c.

2026-06-28 genomics 10.64898/2026.06.23.733953 medRxiv
Top 0.1%
4.9%
Show abstract

BackgroundGray is one of the relatively rare coat colors in donkeys. The Hetian Gray donkey is a distinctive indigenous breed from the Xinjiang Uygur Autonomous Region of Northwestern China, characterized by progressive hair depigmentation with aging while retaining dark skin pigmentation. However, the genetic basis underlying this unique gray coat color phenotype remains unclear. ResultsTo elucidate the genetic basis, we conducted whole-genome resequencing of Gray and non-Gray donkeys. Genome-wide selection signature analyses identified a candidate region on chromosome 15. Subsequent fine-mapping using mass spectrometry-based genotyping of 42 loci refined the candidate interval and revealed a SNP within intron 2 of the ASIP gene, located in a genomic fragment with highly similar sequences, showing complete association with the gray coat color. Association analysis in an expanded population further confirmed a strong correlation between this variant and the gray phenotype. Gene expression analyses also supported the role of ASIP in regulating pigmentation in donkeys. ConclusionsThese findings identify a genetic determinant of gray coat color in donkeys and provide new insights into the molecular mechanisms underlying age-related depigmentation in domestic animals.

9
Variation in AMY2B Copy Number and Serum Amylase Activity in Wolves (Canis Lupus), Brown Bears (Ursus arctos), and Red Foxes (Vulpes vulpes) from Bosnia and Herzegovina

Katica, J.; Crnkic, C.; Kavazovic, A.; Tahirovic, D.; Pojskic, N.; Skapur, V.; Koro - Spahic, A.; Varatanovic, M.; Goletic, T.

2026-07-14 genetics 10.64898/2026.07.09.737415 medRxiv
Top 0.1%
3.2%
Show abstract

The AMY2B gene encodes pancreatic amylase, a critical enzyme for starch digestion. While previous studies have examined AMY2B copy number variation (CNV) in domestic and some wild animals, less is known about wild carnivores inhabiting regions with limited anthropogenic starch exposure. We analyzed blood samples for serum amylase activity and copy number variation in AMY2B gene from 8 wolves (Canis lupus), 11 brown bears (Ursus arctos), and 3 red foxes (Vulpes vulpes) from Bosnia and Herzegovina. AMY2B gene copy number was assessed using droplet digital PCR (ddPCR), and serum amylase activity and glucose levels were quantified. Although the number of fox samples was limited, foxes and wolves consistently harbored two copies of AMY2B, while brown bears exhibited higher CNV (3.67-8.40, mean 5.88). Serum amylase activity was highest in foxes, moderate in wolves, and variable but lower in bears. Despite differences in AMY2B copy number and serum amylase activity, circulating glucose concentrations did not differ significantly among species. Our findings suggest that variation in AMY2B copy number among wild carnivores may be associated with species-specific evolutionary histories and dietary adaptations, providing insight into genomic mechanisms underlying carbohydrate utilization in natural populations.

10
Insights into the genetic architecture of resistance to viral haemorrhagic septicaemia virus in rainbow trout from a genome-wide association study to in vitro CRISPR-Cas9 functional evaluation

Thomas, V.; Collet, B.; Quillet, E.; Marchand, M.; Huetz, F.; Boudinot, P.; Phocas, F.; Lallias, D.

2026-06-11 genetics 10.64898/2026.06.09.731144 medRxiv
Top 0.1%
3.2%
Show abstract

Viral haemorrhagic septicaemia (VHS) is a severe disease affecting rainbow trout (Oncorhynchus mykiss) and a wide range of wild freshwater and marine fish species. VHSV threatens rainbow trout aquaculture, as it may cause 100% mortality in fry. Previous studies identified a quantitative trait locus (QTL) on chromosome 3 associated with resistance to VHSV waterborne challenge and reduced viral replication in fin explants, although these findings were obtained using limited genetic diversity. The objective of this study was to validate and extend the identification of genomic regions associated with resistance to VHSV in the genetically diverse rainbow trout line designated "synthetic." A genome-wide association study (GWAS) was conducted using whole-genome sequences from parents of progeny classified as resistant or susceptible to a VHSV waterborne challenge. While the QTL on chromosome 3 was not validated in the synthetic line, four novel suggestive SNPs associated with survival following VHSV waterborne challenge were identified on chromosomes 6, 8, 17, and 32. Notably, one SNP on chromosome 17 was located within a gene potentially involved in antiviral defence, a paralog of lrp1 (low-density lipoprotein receptor-related protein 1). To further investigate its role, lrp1 function was analysed in vitro using CRISPR-Cas9 genome editing. Three independent lrp1-/- CHSE-EC cell lines were generated and challenged with VHSV. The results showed that lrp1 is not essential for viral entry but may modulate the inflammatory response during VHSV infection in epithelial cell lines.

11
Nitrogen use efficiency in pigs is associated with transcriptomic signatures related to amino acid metabolism, immune activity, and nutrient partitioning

Monney, B.; Ewaoluwagbemiga, E. O.; Kasper, C.

2026-07-01 genomics 10.64898/2026.06.26.733976 medRxiv
Top 0.1%
3.2%
Show abstract

Dietary protein restriction challenges the allocation of amino acids to growth and other physiological functions and therefore requires coordinated metabolic adaptation. Domestic pigs provide an informative system in which to study such responses, because nitrogen retention directly affects lean growth and can be quantified accurately under controlled feeding and housing conditions. Under reduced-protein diets, pigs differ in how effectively they retain nitrogen, and this variation has a genetic basis, making them well suited to investigate the molecular regulation of nitrogen use efficiency (NUE). Here, we characterise differential gene expression and enriched pathways in liver and skeletal muscle of more than 80 pigs with two divergent NUE phenotypes (high and low) maintained under the same protein-reduced, ad libitum dietary conditions. The two NUE phenotypes were clearly distinct at the transcriptomic level, with 177 differentially expressed genes in the liver and 133 in the muscle. In the liver, differential expression and enrichment analyses indicate reduced amino acid catabolism, lower inflammatory and detoxification activity, and a metabolic state that favours lipid processing and insulin-related regulation over the use of amino acids as energy sources. In skeletal muscle, they point to reduced lipid uptake, lower reliance on amino acid oxidation, and a greater emphasis on protein synthesis, translational regulation, mitochondrial energy metabolism, and growth-related processes. These gene-level patterns were supported and extended by pathway and gene-set enrichment analyses. Together, the results suggest that high and low-NUE pigs differ through coordinated, tissue-specific molecular adaptations. Overall, variation in NUE appears to reflect coordinated, tissue-specific differences in how nutrients are allocated between energy use, storage, and lean tissue growth.

12
A gapless Landrace pig genome resolves centromeres and telomeres and highlights telomere repeat structures in different pig breeds

Grove, H.; Stenlokk, K. S. R.; Lien, S.; Gjuvsland, A. B.; Arnyasi, M.; van Son, M.; Kent, M.

2026-06-30 genomics 10.64898/2026.06.25.734473 medRxiv
Top 0.1%
2.6%
Show abstract

Abstract The Duroc-derived reference genome Sscrofa11.1 has provided a critical foundation for pig genomics, providing a high-quality reference genome for accurate variant detection and comparative genomics but does not capture breed-specific variation. Here, we present a near-complete, gap-free genome assembly for the Landrace pig (Landrace_v1, GCA_963921485.1), spanning all 20 chromosomes and totaling 2.6 Gb, including 176 Mb of sequence absent from Sscrofa11.1. Comparative analyses with recently published high-quality pig genomes reveal a conserved centromere organization across breeds, accompanied by substantial variation in repeat composition and length, and identify a pig specific pattern of telomere variant repeats across eight pig breeds. The improved resolution of repetitive regions in Landrace_v1 enables more complete reconstruction of complex gene families, including olfactory receptors, and uncovers structural variation at the KIT proto-oncogene receptor tyrosine kinase locus not represented in the Duroc reference. Together, these findings highlight the limitations of single-reference genomes and demonstrate the value of breed-specific assemblies for capturing genomic diversity and improving downstream analyses.

13
Genome-wide meQTL mapping in cattle blood reveals cis and trans regulation of DNA methylation

Fouere, C.; Costes, V.; Besnard, F.; Le Danvic, C.; Patry, C.; Fritz, S.; Boussaha, M.; Jouin, M.; Boichard, D.; Kiefer, H.; Costa Monteiro Moreira, G.; Sanchez, M.-P.

2026-07-08 genetics 10.64898/2026.07.07.736355 medRxiv
Top 0.1%
2.4%
Show abstract

Background Complex traits are influenced by numerous variants, most of which have regulatory effects on gene expression that can be mediated by DNA methylation. Molecular QTL mapping is an approach that aims to dissect these effects. However, obtaining molecular phenotypes on a large scale is challenging, particularly in livestock species. In cattle, an epigenotyping array called EpiChip has recently been developed in the European RUMIGEN project. The EpiChip, which contains 43,317 CpG sites distributed all over the bovine genome, enables large-scale measurement of DNA methylation. This study aims to characterize the genetic determinism of blood DNA methylation in cows by estimating heritability and mapping cis- and trans-methylation QTLs (meQTLs). Results Whole blood samples from 4,457 genotyped Holstein cows were epigenotyped. Across all CpG sites, the heritability estimates averaged 24.6%. The local meQTL mapping at sequence-level for variable CpG sites (SD > 2.5%; n = 28,806) detected cis-meQTLs for 80.1% of the CpG sites, with sentinel SNPs located close to their associated CpGs. A two-step analysis was also conducted to identify long-range associations, with a particular focus on trans-meQTL hotspots. First, we identified CpG-SNP trans-associations using medium-density genotypes (50k SNPs) that revealed 31,846 SNPs with significant effects on 1 to 530 trans-CpG sites. Then, regions associated with at least 34 independent trans-CpGs were retained defining 31 hotpots. For each hotspot, a local sequence-level GWAS was conducted using the first principal component derived from the associated trans-CpGs. Out of the 31 detected hotspots, three were located close to transcription factor genes (RUNX1, NFIC and FOXA3) for which the associated trans-CpGs were enriched for the corresponding binding motif. Two other hotspots were located within KDM5A and KDM5B, and their corresponding trans-CpGs were strongly overrepresented in H3K4me3 narrow peaks in blood as well as in other tissues. Conclusions By identifying functional candidate genes associated with blood DNA methylation in cattle, these findings provide new insights into the regulatory architecture of DNA methylation in mammals, highlighting the value of large-scale molecular data from livestock populations.

14
Vgll3a promotes sexual maturation in male and female Atlantic salmon

Kjaerner-Semb, E.; Fraser, T. W. K.; Vogelsang, P.; Skaftnesmo, K.; Ayllon, F.; Edvardsen, R. B.; Braathen, S.; Norberg, B.; Fjelldal, P. G.; Andersson, E.; Schulz, R. W.; Wargelius, A.

2026-06-19 genomics 10.64898/2026.06.19.733363 medRxiv
Top 0.1%
2.3%
Show abstract

The age at which Atlantic salmon reaches sexual maturity shows a strong hereditary component associated with the vgll3a locus. The role of Vgll3 in maturation has remained unknown in vertebrates until recently, when it has been linked to pleiotropic roles in killifish, both delaying male maturation and affecting lifespan by protecting against cancer. As Atlantic salmon has two vgll3 paralogs, where only vgll3a has been associated with sexual maturation, it may provide a suitable model for studying the maturation-specific function of vgll3, as the other paralog may buffer for pleiotropic roles of vgll3. To address this, we used CRISPR/Cas9 to generate fish highly mutated in the vgll3a gene. We monitored their maturation and crossed highly mutated crispants to generate two year-classes of complete loss-of-function. All groups were reared under environmental conditions triggering early maturation in one-year-old males. We found a clear difference in the proportion of sexually maturing or mature fish between the different genotypes: in all experiments significantly fewer vgll3a-/- males entered puberty and reached final maturation compared to vgll3a+/- and vgll3a+/+ males. Furthermore, loss of vgll3a resulted in lower frequencies of maturation also in females. We conclude that Vgll3a stimulates maturation and that its complete removal significantly reduced maturation rates in both sexes in Atlantic salmon. Our findings also identify vgll3a as the causative gene in the locus associated with age at sexual maturity. Together, our findings support a new role for Vgll3 in initiating puberty in vertebrates and identifying salmon as a promising model for functional studies regarding the timing of sexual maturation.

15
Pangenome Graph Node-Phenotype Association shows GWAS-like quality results with only few individuals

Carrette, C.; Sabot, F.; Muller, C.

2026-08-01 bioinformatics 10.64898/2026.07.31.741971 medRxiv
Top 0.1%
2.1%
Show abstract

PurposeWe introduce GO_SCPLOWRAC_SCPLOWNPA, standing for Graph Node-Phenotype Association, a method performing a GWAS-like analysis on a pangenome variation graph (PVG) built using a small number of individual genome sequences, without the need for additional population materials or kinship information for qualitative phenotypes. This method reduces the number of individuals required for association studies and prevents reference bias from variant calling in these types of analyses. BackgroundA PVG represents the multiple alignment of a set of complete genomes. It contains all variations, from single nucleotide polymorphisms (SNPs) to large structural variations (SVs), which are represented as nodes in the graph. By integrating phenotype information within nodes, we can assign a Phenotype Score (PS) to each node in the PVG and identify phenotype-related regions directly within it. These regions represent statistically significant shifts in PS distribution, highlighting their implication in the phenotype. Finally, GO_SCPLOWRAC_SCPLOWNPA provides their positions and scores for further analysis. ResultsThis method was tested using simulated data and two publicly available datasets: the Sub1A gene locus for Oryza sativa in a 13 individuals PVG, and the insertion responsible for the white-headed cattle with a PVG of 24 individuals. Source code of GO_SCPLOWRAC_SCPLOWNPA is available here https://forge.ird.fr/diade/graphgwas/granpa under GNU GPLv3. ConclusionGO_SCPLOWRAC_SCPLOWNPA was able to identify the expected area in two simulated datasets and the responsible loci for these two known traits using only a few dozen complete genomes in these PVGs. While currently limited to qualitative phenotypes, this method opens the way to more efficient ones relying on PVGs and few individuals.

16
Bridging Morphology and Genomics: A rapid image-based assessment of genomic admixture in the endangered gayal (Bos frontalis)

Ma, J.; Chen, Y.; Guo, Z.; Xiao, J.; Wu, H.; Luo, J.; Zhang, Y.-p.; Li, Y.

2026-08-25 zoology 10.64898/2026.08.25.746947 medRxiv
Top 0.1%
2.0%
Show abstract

Abstract The gayal (Bos frontalis) is an endangered semi-domesticated bovine species renowned for its high-quality beef. However, its semi-feral lifestyle, ongoing habitat fragmentation, and extensive genetic introgression from sympatric local cattle have led to dramatic population decline and severe erosion of purebred genetic integrity, posing substantial challenges to its conservation and utilization. To address the urgent demand for rapid, non-invasive, and field-compatible germplasm identification, we developed an integrated artificial intelligence (AI) framework that predicts genomic admixture composition from external morphological images. We constructed a comprehensive dataset comprising 6,245 morphological images and matched genomic sequences from 52 gayals maintained at the Yunnan Provincial Gayal Conservation Farms. Following a preliminary evaluation of nine deep learning models, five were incorporated into a anatomical segment-based multi-modal pipeline, among which Inception_V3 delivered the optimal overall performance. To enhance simultaneous extraction of local fine-grained features and global structural information, we further designed an innovative HybridInceptionViT model by integrating the multi-scale Inception module with the Vision Transformer (ViT) framework. This hybrid model significantly outperformed the baseline Inception_V3, boosting the accuracy of phenotype-derived prediction against genomic admixture estimate from 69.69% to 87.87% (absolute error <15%). This study establishes a practical, low-cost "phenotype-to-genotype" tool for rapid on-site gayal germplasm screening, offering a scalable strategy for the conservation and breeding management of endangered livestock, and holds broad application prospects for agricultural and livestock production systems.

17
Modeling population control via tunable sex ratio distorter gene drives in Aedes aegypti

Childs, L. M.; Shabani, S.; Tauber, U.; Tu, Z.

2026-07-09 genetics 10.64898/2026.07.05.736587 medRxiv
Top 0.1%
1.7%
Show abstract

Aedes aegypti is a major vector of arboviruses, and belongs to subfamily Culicinae, a diverse group of mosquitoes with homomorphic sex-determining chromosomes. Males are the heterogametic sex with a dominant male-determining locus (M locus). The M locus and its counterpart m locus are embedded in a region of suppressed recombination, with a large portion of this recombination desert showing significant molecular differentiation despite homomorphy. We developed a mathematical framework to examine M-linked genome editors that specifically target the m-chromosome during spermatogenesis, mimicking the naturally occurring sex ratio distorters (SRDs) in Culicinae that produce male-biased meiotic drives. Unlike previous models for species with heteromorphic sex chromosomes (e.g., X and Y), we incorporate features stemming from the homomorphic nature of the Ae. aegypti sex chromosomes such as varied linkage to the M locus, making the degree of super-Mendelian inheritance readily tunable. We evaluated in silico SRDs with a range of M-linkage and editing efficiencies and established the theoretical foundation for developing highly efficient SRDs that outperform several methods of population suppression. These SRDs can be tuned to mitigate impact on a neighboring population. The framework developed here is suitable for exploring SRD-mediated genetic biocontrol of pests with homomorphic sex chromosomes.

18
Comparison of localGEBV and Optimal Haplotype Stacking Fitness Functions using a Novel R Package: HapSelect

Shaffer, W.; Papin, V.; Carter, Z.; Brunner, S. M.; Tong, J.; Villiers, K.; Robinson, H.; Voss-Fels, K.; Hayes, B. J.; Hickey, L.; Dinglasan, E.

2026-07-13 genetics 10.64898/2026.07.08.737160 medRxiv
Top 0.1%
1.7%
Show abstract

Haplotype-based breeding strategies have emerged as promising approaches to maximize long-term genetic gain by identifying complementary parental combinations while maintaining genetic diversity. However, these methods typically require phased genotypes and more intensive workflow pipelines and skillsets. We developed a novel local genomic estimated breeding value (localGEBV) fitness function with similar intent to the optimal haplotype stacking (OHS) framework fitness function and implemented both in the novel R package, HapSelect. Our aim was to evaluate whether phased haplotypes provide additional benefit over the more easily available dosage-based unphased genotypes in highly inbred crops. A subset of bread wheat nested association mapping (NAM) population comprising 444 lines genotyped with 6,054 DArT-Seq markers was analysed. Marker effects were estimated using rrBLUP, localGEBV and haplotype effects were calculated across linkage disequilibrium-defined haploblocks, and genetic algorithms (GA) were used to identify optimal sets of 30 founders using either a localGEBV derived fitness function with unphased, dosage inputs or the OHS fitness function with phased inputs. Selected parental sets were compared with conventional truncation selection (TS) through 150 generations of forward simulation. The OHS fitness function achieved a marginally greater optimized ultimate GEBV than the localGEBV fitness function during GA optimization, with only 18 of the 30 selected founders overlapped between the two methods. Despite these differences, forward simulations demonstrated nearly identical long-term genetic gain for localGEBV and OHS-selected founders, with both approaches outperforming conventional truncation selection by maintaining greater genetic diversity and delaying the genetic plateau. The minimal difference between localGEBV and OHS is likely attributable to the high homozygosity of the population, where localGEBV and haplotype effects are nearly confounded. These results demonstrate that dosage-based localGEBV provides a practical alternative to phased haplotype approaches for parent selection in inbred crops, substantially simplifying genomic workflows while maintaining long-term breeding performance. Future work should evaluate these methods in more diverse inbred populations and outbred species, where great haplotypic diversity may increase the advantage of true haplotype-based optimizations.

19
Haplotype assembly without parental sequencing: Genotype-based trio-binning (GT-Trio)

Hettasch, T. J.; Gjuvsland, A. B.; Kent, M. P.; Grove, H.; Vage, D. I.

2026-06-11 genomics 10.64898/2026.06.08.729486 medRxiv
Top 0.1%
1.7%
Show abstract

Trio-binning is a robust method for haplotype-resolved assembly, providing the most accurate representation of diploid genomes including complex and haplotype-specific variation. Conventional trio-binning methods depend on parental short-read sequences to differentiate offspring reads originating from the maternal and paternal haplotypes. Here, we present a genotype-based trio-binning pipeline (GT-Trio) which reconstructs parent sequences from phased parental genotypes and uses this as an alternative source of parental information for haplotype assembly. The GT-Trio pipeline was applied to assemble the maternal and paternal haplotypes of three Norwegian Red (NR) cattle individuals, using phased parental genotypes imputed from array to sequence as input. Haplotypes assembled with GT-Trio using all sequence variants as parental input demonstrated assembly quality and phasing accuracy comparable to that achieved with conventional trio-binning. Using lower density subsets of array SNPs led to a slight reduction in accuracy of haplotype separation, accompanied by an increase in size, contiguity and completeness, suggesting a trade-off between assembly quality and phasing accuracy associated with the density of parental genotypes provided as input to the pipeline. Overall, GT-Trio provides a scalable framework for haplotype assembly without parental sequencing and will be applicable in livestock species where genotyping and imputation is performed routinely. The GT-Trio pipeline is available at https://github.com/theahettasch/GT-Trio.

20
Cross Potential Selection for Multiple Traits Considering the Progeny Distribution of Future Inbred Lines in Plant Breeding Programs

Sakurai, K.; Moreau, L.; Mary-Huard, T.; Charcosset, A.; Iwata, H.

2026-06-08 genetics 10.64898/2026.06.02.729654 medRxiv
Top 0.1%
1.6%
Show abstract

In plant breeding, it is often necessary to improve a target trait while maintaining other essential traits within desirable ranges. When genetic relationships exist among these traits, improvements in the target trait may lead to undesirable changes in essential traits, complicating cross selections. In such cases, it is critical to select cross-pairs that are expected to produce progeny that satisfy the requirements for all traits. The progeny distribution of each crossing pair can be predicted using the estimated genotypic values and genetic (co)variances of the target and essential traits. By utilizing this distribution, the probability of generating progeny that satisfy predefined trait requirements can be evaluated, allowing a direct comparison of alternative crosses. In this study, we developed Cross Potential Selection for Multiple Traits (CPS-MT), a breeding strategy designed to improve a target trait while maintaining one or more essential traits within desirable ranges. CPS-MT extends the original Cross Potential Selection (CPS) framework to explicitly handle trade-offs between traits under genetic correlations. We evaluated the performance of CPS-MT through simulations involving four types of genetic relationships and two genetic causal factors between traits, resulting in seven scenarios. Across all scenarios, CPS-MT consistently improved the likelihood of obtaining desirable progeny, indicating that CPS-MT provides a practical and effective framework for cross selection under multi-trait constraints in breeding programs. Article SummaryThis study developed Cross Potential Selection for Multiple Traits (CPS-MT), a new breeding strategy designed to improve a target trait while maintaining one or more essential traits within desirable ranges. CPS-MT evaluates crossing pairs by predicting progeny distributions based on estimated genotypic values and genetic covariances, enabling direct comparison of alternative crosses under multi-trait constraints. Through simulations incorporating four types of genetic relationships and two causal factors (seven scenarios), CPS-MT consistently increased the likelihood of obtaining progeny that satisfied the predefined trait requirement. These results indicate that CPS-MT provides a practical, robust framework for target trait improvement under trait constraints.